Statistical Methods in Medical Research
○ SAGE Publications
Preprints posted in the last 90 days, ranked by how well they match Statistical Methods in Medical Research's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
MA, Z.; XIANG, Y.; So, H.-C.
Show abstract
Abstract Purpose This study introduces a novel approach to address unmeasured confounding in terminal event studies using the prior event rate ratio (PERR) method. The proposed approach PERR_{proxy} used a proxy event to replace the original terminal event in the pre-exposure period, enabling the application of PERR in terminal event settings. Additionally, we also applied difference in difference (DID) regression, which is conceptually analogous to PERR to estimate the standard errors and confidence intervals of PERR_{proxy}. Methods We conducted numeric simulations to evaluate the validity of PERR_{proxy} approach and assessed its performance under varying levels of unmeasured confounding effects, baseline hazard ratios, and the correlation between the proxy and terminal events. To demonstrate its practical applicability, we also performed an empirical analysis to investigate the impact of severe hospitalized COVID-19 on circulatory system disease mortality using the PERR_{proxy}. Results In simulation studies, PERR_{proxy} effectively reduced the unmeasured confounding effects compared to the conventional methods. The performance of PERR_{proxy} was influenced by the strength of unmeasured confounding, baseline hazard ratios, and the correlation between the proxy and terminal outcomes. In addition, difference in difference (DID) regression had much faster computational speed for estimating standard errors and confidence intervals compared to bootstrap. In the empirical analysis, PERR_{proxy} identified that severe hospitalized COVID-19 as a significant risk factor for the circulatory system disease mortality and reduced the unmeasured confounding effects. Conclusions The PERR_{proxy} approach extends the applicability of the original PERR method to terminal event studies, offering a promising solution for addressing unmeasured confounding. Additionally, the DID regression framework provides a computationally efficient alternative for parameter estimation in PERR-based studies. However, careful consideration is still required in PERR_{proxy} for proxy events selection and other underlying assumptions of the PERR method to ensure valid results. Keywords: prior event rate ratio, unmeasured confounding, proxy event, terminal event study, observational study, electronic health records
Mell, L. K.
Show abstract
In competing risks settings, covariate effects and group comparisons are usually assessed one event at a time - through log-rank or Cox tests on the cause-specific hazards, or Gray's test or Fine-Gray regression on a cumulative incidence function (CIF). This can obscure a clinically important quantity: the ratio between the event of interest and the competing event, since groups may differ little on the individual events yet differ sharply in their ratio. The generalized competing event (GCE) framework makes this ratio the object of inference; on the cause-specific scale the hazard ratio omega+(t) = lambda_1(t)/lambda_2(t) is estimated efficiently from a single stacked (Lunn-McNeil) model. We extend the framework to two scales that describe realized incidence. The subdistribution hazard ratio omega-tilde+(t) = lambda-tilde_1(t)/lambda-tilde_2(t) is estimated by a stacked, risk-set-weighted extension of the Lunn-McNeil construction; the cumulative-incidence ratio rho(t) = F_1(t)/F_2(t) - the odds that a subject's realized event by time t is the event of interest - by jackknife pseudo-observation regression of the Aalen-Johansen estimator. We relate the three contrasts: rho equals omega+ exactly under proportional cause-specific hazards, and equals omega-tilde+ only in the small-time limit under proportional subdistribution hazards, drifting toward 1 thereafter. The orthogonality that makes omega+ efficient is lost on both cumulative-incidence scales - omega tilde+ through overlapping weighted risk sets and shared censoring weights, rho through the shared all-cause survivor - so each carries a covariance term that must be handled and that bounds efficiency relative to the hazard-scale test. We derive the corresponding variances, study operating characteristics by simulation, illustrate on hypothetical prostate and head-and-neck cohorts, and provide an implementation in the gcemod R package.
Qian, Y.; Song, Y.
Show abstract
Instrumental variable (IV) methods are widely used in health and social sciences to estimate causal treatment effects among compliers. In certain research settings, the instrument-treatment association (first stage) and the instrument-outcome association (reduced form) are each estimated from a different dataset. Two-Sample Instrumental Variables (TSIV), proposed by Angrist and Krueger (1992), addresses this by combining first-stage and reduced-form estimates from separate data sources into a single causal effect estimate. However, TSIV identification requires that instrument compliance behavior be consistent across the two samples, a condition that is rarely verified in practice. We show mathematically and empirically that when compliance differs between samples, the raw TSIV estimator does not converge to the true Local Average Treatment Effect (LATE) and instead attenuates toward a predictably biased limit proportional to the ratio of first-stage compliance rates between the two samples. To address this, we formalize a framework for estimating LATE with TSIV under two key assumptions: (1) Covariate Overlap, requiring that the two samples share sufficient common support in their covariate distributions, and (2) Compliance Transportability, requiring that compliance behavior is identical across populations after conditioning on observed covariates. We consider a setting in which a health policy instrument and outcomes are recorded in administrative claims while treatment and covariates are collected in a survey. We use a C-statistic derived from pooled covariates to detect population mismatch and an Inverse Probability Weighting (IPW) correction that reweights the first-stage sample to approximate the administrative covariate distribution. In Monte Carlo simulations across eight scenarios calibrated to a survey-Medicaid setting, IPW-TSIV reduces bias in estimating the LATE, achieving 88% reduction in the primary scenario, 82% under severe selection, and 79% when state-level expansion policy drives compliance heterogeneity. We further validate this framework using the Oregon Health Insurance Experiment, where partitioning the public-use lottery data (N = 24,646) into two non-overlapping samples with substantively meaningful compliance heterogeneity yields a verifiable benchmark against the true causal effect. IPW-TSIV reduces mean absolute bias by 71.6% relative to the oracle S2-specific LATE across 10 independent replications (C-statistic = 0.78), outperforms naive TSIV in all 10 splits, and reduces mean bias relative to the full-data LATE from +0.016 to +0.008. This framework provides applied researchers with actionable diagnostic thresholds to detect sample mismatch, validate transportability assumptions, and determine when structural TSIV estimation is reliable.
Jah, A.; Ngesa, O.; Wamwea, C.; Ngunyi, A.
Show abstract
Abstract: Compartmental epidemic models conventionally treat the probability of moving between dis ease states as fixed over time, an assumption that sits uneasily with the reality of a pandemic in which lockdowns, mask mandates, vaccination roll-out, and the arrival of new variants continually reshape transmission. This paper develops a time-inhomogeneous Markov chain framework for the Susceptible-Exposed-Infectious-Removed (SEIR) process, in which each transition probability pab(t) is allowed to vary with calendar time while respecting the struc tural zeros implied by the SEIR compartmental flow. We derive the constrained maximum likelihood estimator of pab(t) under these structural constraints, establish its finite sample efficiency, asymptotic normality, and Wilson score confidence intervals, and construct a like lihood ratio test of the null hypothesis that a compartments exit probability is constant over time. We further propose a stochastic machine learning hybrid extension in which the raw, kernel smoothed transition probabilities are regressed on policy and mobility covariates using both a logistic generalized linear model and a random forest, allowing the framework to attribute time-inhomogeneity to observable interventions. The methodology is applied to a compiled daily, district level COVID-19 surveillance panel for Sierra Leone spanning March 2020 to December 2023 (16 districts, 1,401 days). The likelihood ratio test rejects time-homogeneity of the exposed to infectious transition in 15 of 16 districts and of the infectious-to-removed transition in 8 of 16 districts ( = 0.05), and the covariate augmented logistic model achieves an out of sample Brier score roughly 76 times smaller than a time homogeneous pooled baseline, with healthcare capacity and the time trend emerging as the most influential predictors in the random-forest component. These results provide statisti cal evidence that time-inhomogeneous, covariate informed Markov models offer a materially better description of district-level COVID-19 transmission in Sierra Leone than classical time homogeneous compartmental models, with implications for sub-national outbreak monitoring in resource constrained settings
Schwenke, J. M.; Herkner, F.; Kayembe, M. T.; Olsen, I. C.; Briel, M.; König, F.
Show abstract
Acute viral respiratory infections (ARVIs) are a major cause of hospitalization and death worldwide, yet randomized clinical trials in this setting face substantial challenges in selecting efficient and clinically meaningful primary endpoints. Mortality is often too infrequent to serve as a feasible primary endpoint. Several alternative approaches have been proposed, including ordinal scales, time-to-event endpoints, recovery-based composite outcomes, and longitudinal ordinal models. However, their comparative operating characteristics under realistic ARVI disease courses remain insufficiently understood. We describe a simulation study to compare the type I error and power of commonly used and recently proposed endpoints and analysis strategies for two-arm randomized trials in hospitalized participants with ARVIs. Data will be generated under several mechanisms designed to mimic plausible participant trajectories, including a latent Brownian motion process, a first-order ordinal Markov process, a latent recurrent-event process with frailty, and resampling from individual participant data from the ACTT-2 trial. Simulated outcomes will use 4-, 6-, and 8-level ordinal severity scales and will reflect moderately and severely ill populations, follow-up horizons of 28 or 60 days, varying treatment effects, and sample sizes. Methods to be compared include Markov ordinal state transition models, proportional-odds models at a fixed time point, days-to-recovery scale analyses, Cox models for time-to-event endpoints, logistic regression for binary endpoints, generalized pairwise comparisons for hierarchical composites, and t-tests for days alive and out of hospital. This study will provide a systematic comparison of endpoint definitions and analysis methods for ARVI trials under clinically motivated data-generating mechanisms. The results are intended to inform the selection of feasible, interpretable, and statistically efficient primary analysis strategies for future trials in viral respiratory disease.
Ng, S.-P.
Show abstract
The incidence rate ratio R is the standard measure for comparing event rates in clinical trials and epidemiology. In vaccine trials, the vaccine efficacy is VE = 1 - R. When events are rare, the two arm counts are Poisson. The estimator of R is heteroskedastic: its sampling variance changes with the data. So no fixed-width interval covers correctly everywhere. The usual log-Wald interval is undefined at zero events and covers poorly at small counts. Early vaccine and drug-safety readouts fall in exactly this regime. We show that a single reparameterization collapses this bivariate problem to an effective one-parameter family with a quadratic variance function, whose variance-stabilizing transformation is 2 arcsinh(sqrt(R)). The reduction yields a closed-form confidence interval for R. Its two leading errors, a curvature bias and the variability of the estimated scale, each admit a closed-form correction with no tuning constants. In a Monte Carlo study of our seven arcsinh variants and five competitors, the +Curve+Stu variant covers within 0.002 of the nominal 0.95 for about 50 control and 5 treatment events. Its width is on par with the best competitor. It avoids the conservatism and zero-count breakdown of log-Wald and MOVER. For moderate counts, we recommend this interval; for sparser data, our Bar-Lev and Enis count-shift variant is more robust. The result is a ready-to-use, closed-form interval for the low-count regime. We illustrate it on early Covid-19 vaccine-efficacy readouts and provide reference implementations in R and Python.
Tzamarias, B. D. E.; Burroughs, N.
Show abstract
Cancer therapy balances between two competing objectives - treatment efficacy against the tumour and the risk of treatment related severe adverse events, including patient death. Most existing optimal control theory (OCT) formulations rely on optimising heuristic cost functionals that lack direct clinical interpretability. In clinical practice treatment efficacy and patient tolerability are primarily assessed through survival metrics and adverse event rates. Here we introduce the Continuous Lifetime Payoff (CLP), a novel OCT objective functional that directly links treatment decisions to patient survival. It explicitly incorporates tumour dynamics, tumour eradication, and patient mortality from tumour progression, drug-related toxicity and age. We fit age-related mortality from life tables and infer parameters from simulated survival data. The CLP provides a clinically grounded framework for optimising chemotherapy regimens.
Wang, H.; Zhang, B.; Lei, Y.; Lu, Y.; Zhang, D.; Jian, X.; Zhu, Y.; Hu, W.; Chu, H.; Chen, Y.; Suchard, M. A.; Ryan, P. B.; Hripcsak, G.; Asch, D. A.; Lu, Y.; Bin, Y.; Schuemie, M. J.; Qiu, Y.; Chen, Y.
Show abstract
Glucagon-like peptide-1 receptor agonists (GLP-1RAs) have been linked to heterogeneous, potentially pleiotropic effects across organ systems, motivating outcome-wide comparative risk profiling in real-world data. A central challenge in such analyses is \emph{residual bias} that remains after adjustment for observed confounders, which can distort effect estimates and mis-calibrate uncertainty. We present distributional diagnosis and calibration (DC), which uses panels of negative control outcomes (NCOs) to diagnose residual bias and calibrate uncertainty. DC evaluates null behavior via $p$-value uniformity and empirical coverage across NCOs, and uses the empirical distribution of NCO effect estimates to calibrate confidence intervals for prespecified primary outcomes. DC is modular: it can wrap around commonly used causal inference methods and operates directly on summary statistics, supporting collaborative research under data-sharing constraints. Using electronic health records from a large U.S. clinical research network (152.7 million patients), we compared GLP-1RAs with sodium--glucose cotransporter~2 inhibitors across 15 prespecified outcomes spanning cardiovascular, mental health, and genitourinary domains using four causal estimators. Across outcomes and methods, DC diagnostics revealed substantial and method-dependent residual systematic error. DC calibration attenuated systematic error signals observed in negative controls and yielded more stable, better-calibrated estimates for clinical outcomes, supporting DC as a practical strategy to strengthen the credibility of real-world comparative effectiveness research.
Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.
Show abstract
Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.
Otte, W. M.
Show abstract
Meta-analysis usually reduces each study to an effect estimate with a standard error and pools these by inverse-variance weighting: fixed effect (FE), random effects (RE), or unrestricted weighted least squares (UWLS). We propose information-geometric meta-integration (IGMI), representing each study by its sampling distribution, the Gaussian N(theta_i, Sigma_i), and pooling studies as a weighted Frechet mean (barycenter) under Bures-Wasserstein (BW), Fisher-Rao, or Wasserstein-Fisher-Rao (WFR) geometry. In the scalar fixed-variance case the BW barycenter mean is exactly the FE estimate; the minimized Frechet functional reproduces the Higgins-Thompson I^2 and DerSimonian-Laird tau^2 heterogeneity statistics; and a Frechet-scatter pivot reproduces the Hartung-Knapp-Sidik-Jonkman interval at m = 1 and yields an exact Hotelling F(m, K-m) region for m outcomes under proportional total covariances. WFR adds a robust outlier-resistant pool: as its length scale delta grows without bound it converges monotonically to BW, whereas finite delta gives a redescending M-estimator with rejection point exactly pi*delta. Simulations show calibrated multivariate coverage at small K, where Wald intervals undercover, and strong resistance of the equal-weight WFR pool to contamination. In 2,445 Cochrane meta-analyses, WFR most often wins leave-one-out predictive scoring. In 835 bivariate meta-analyses, the closed-form BW barycenter matches REML multivariate meta-analysis predictively and is exactly invariant to the unreported within-study correlation, unlike the likelihood estimate.
Di Carluccio, E.; Koliopanos, G.; Ojeda, F. M.; Weimar, C.; Ziegler, A.
Show abstract
Statistical prediction models for binary outcomes are becoming increasingly popular. One significant challenge is calibrating these models to suit the characteristics of a target population that is structurally different from the original population. Calibration is especially challenging when there is no training data available from the target population. To address this problem, we propose a novel calibration method, SimCal, which uses synthetic data generated from the model development data in conjunction with marginal statistics from the calibration cohort. We show that expert judgment modeling (EJM) may be used for calibration if cross-sectional data from the target population are available comprising expert judgments about the potential outcome and the covariates. We describe three alternative calibration approaches when calibration data are lacking: similarity-binning averaging (SBA), adaptive calibration of predictions (ACP), and Elkan calibration. In a simulation study, we compare SBA, ACP, Elkan calibration, and SimCal. R code for applying these methods is provided from the re-analysis of data on coronary artery disease. We illustrate all 5 calibration approaches with a real data set for predicting functional outcome after stroke and all approaches but EJM in the re-analysis of the Cleveland Clinic data. None of the approaches performed convincingly well in all situations. SimCal performed well when model parameters were correctly specified. EJM failed on the stroke data. Further research is urgently required for calibration in the absence of calibration data.
Yang, F.; Magee, A.; Morris, S. E.; Mathis, S. M.; Wiegand, R.; Iuliano, D. A.; Biggerstaff, M.; Olesen, S. W.
Show abstract
Vaccination can be a useful intervention for reducing infectious disease burden. Estimating numbers of vaccine-prevented health outcomes is one approach to quantifying the benefits of vaccination. Here we improve a method described by Foppa et al. (1) that assumes vaccination has only direct effects, that is, it cannot prevent infection or onward transmission of the disease. We rederive this method and derive an improved method that increases estimation accuracy with minimal additional analytical complexity. To evaluate the improved method, we simulated disease outbreaks and compared the accuracy of the two methods for estimating prevented disease outcomes. In 84% of simulations performed over a wide parameter space, the improved method had an equal or smaller estimation error compared to the original Foppa method, with 7.9-fold smaller mean error and 44-fold smaller standard deviation of errors. Our study improves a method for estimating prevented burden when assuming vaccination has only direct effects.
Ahmad, A. K.; Pandrich, M.; Naik, A.; Astruc, A.; Lafferty, K.; Shah, N. M.; Ofili-Yebovi, D.
Show abstract
Background: Early access to pregnancy assessment units now detects many tubal ectopic pregnancies (TEP) at a stage when they could resolve spontaneously, creating a management dilemma. Methods: We performed a hypothesis-generating exploratory analysis in a retrospective study to assess whether serum progesterone (P4) levels in women with TEP are associated with management outcome. Results: Ninety-one cases of TEP managed in a single centre over three years were analysed. Receiver operating characteristic (ROC) curve analysis was used to explore serum levels of progesterone (P4), first human chorionic gonadotropin (hCG) and peak hCG (alone and in combination) in relation with successful completion of expectant management. Decision-tree analysis using first hCG and P4 was additionally performed to explore clinical sequential risk stratification. 23% (n=21) successfully completed expectant management. P4 concentrations in the expectant management group (median 3 nmol/L, IQR 2.00 to 8.50) were significantly lower than in those requiring surgical or medical management (median 17 nmol/L, IQR 5.75 to 29.25; p=0.0002). Area under the ROC curve (AUC) values for P4, log10 first hCG, log10 peak hCG and P4 with log10 first hCG were 0.766, 0.814, 0.811 and 0.835, respectively, for predicting successful expectant management. However, hCG was not significantly outperformed. Nonetheless, Youden optimised thresholds for hCG and P4 are reported, alongside decision-tree analysis that identified sequential first hCG and P4 thresholds associated with successful expectant management. Conclusion: Lower P4 levels are associated with successful expectant management of TEP but they do not outperform hCG either alone or as an adjunctive marker.
Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.
Show abstract
Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.
Das, N.; Ueki, M.
Show abstract
Population stratification is a major source of inflated false positive rates in genome wide association studies. However, relatively few studies have examined its impact on gene-gene interaction detection, despite the importance of epistasis for understanding the genetic architecture of complex traits. In this study, we identify scenarios under which population stratification can inflate the interaction test statistics. Through analytical derivations and simulation studies, we show that this inflation is not adequately controlled by including principal components as covariates in the regression model. We then propose an alternative approach that effectively controls the inflation of false-positive rates for interaction test statistics due to population stratification by using single nucleotide polymorphism-by-population structure interaction as an additional covariate term in the regression model.
Chen, T.; Voorhies, K.; Reeson, A.; Seo, S.; Lee, S.; Hahn, G.; Hecker, J.; Prokopenko, D.; Hoth, K.; Kelly, R.; Lasky-Su, J. A.; Weiss, S.; Lange, C.; Lutz, S.
Show abstract
Mendelian Randomization (MR) is a popular tool for inferring causal relationships between traits using genetic variants as instrumental variables. These methods have been extended to also determine the direction of causality. However, causal direction cannot be inferred from a statistical test or estimation procedure (i.e. from data alone) without further assumptions and the methods operating characteristics and relative performances are not well understood. We conducted a comprehensive simulation study to illustrate this issue by evaluating type I error and power of 17 summary-based MR methods for inferring the effect direction. These methods fall within three methodological families: MR Steiger, Causal Direction (CD), and bidirectional MR approaches, with scenarios ranging across combinations of horizontal pleiotropy, unmeasured confounding, measurement error, longitudinal feedback, and varying sample sizes. While most methods achieved sufficient power levels under the alternative hypothesis in most scenarios, we found that every method was susceptible to inferring the wrong causal direction or under powered, and no method consistently maintained both correct type 1 error control and high power. In our applications, we evaluated the effect direction between the trait pairs body mass index (BMI) and major depressive disorder (MDD) and between BMI and asthma. To help researchers to evaluate the 17 methods to infer the effect direction and consider these challenges in their own data, we have developed MRdirection, an R package that runs the simulation studies examining the 17 directional MR methods across different user-defined scenarios. Our study, together with the accompanying R package, provides researchers with a tool for examining directional MR methods given different underlying assumptions.
ZHAO, M.; LIU, J.; HAN, D.; ZHANG, C.; ZHOU, Y.; CHEN, S.; LIU, C.
Show abstract
Objective: To perform internal and external validation of a gradient-boosted decision tree (GBDT) fusion model that integrates zygote morphokinetic parameters with conventional embryo assessment features for blastocyst prediction, and to compare its discriminative performance against senior embryologists. Methods: This retrospective cohort study included 631 normally fertilized zygotes from 218 treatment cycles. A GBDT fusion model integrating 84 zygote morphokinetic parameters and 8 conventional assessment features was evaluated internally (5-fold cross-validation) and externally on a public dataset of 523 embryos with blastocyst outcomes. Model performance was assessed using area under the ROC curve (AUC), area under the precision-recall curve (AUPRC), F1 score, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Discrimination was compared with embryologist consensus using the DeLong test; agreement was assessed with Cohen's kappa. Results: The model achieved an internal AUC of 0.78 (95% CI 0.74-0.82), AUPRC 0.72, F1 0.73, sensitivity 0.74, specificity 0.77, PPV 0.72, and NPV 0.79. External validation on the public dataset demonstrated acceptable generalizability (AUC 0.76, 95% CI 0.71-0.81). The model significantly outperformed embryologist consensus (AUC 0.70, P<0.001) with moderate agreement (kappa=0.56). Decision curve analysis confirmed clinical net benefit at threshold probabilities of 0.15-0.55. Conclusions: The GBDT fusion model integrating zygote morphokinetics with conventional assessment demonstrates good discrimination and external generalizability for blastocyst prediction, providing an interpretable decision-support tool for embryo selection in IVF practice.
Hsu, C.-Y.; Liu, Q.; Shyr, Y.
Show abstract
As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.
Pan, W.; Lu, Z.; Jiang, W.; Lim, J.; Xu, L.; Wang, X.
Show abstract
In meta-analyses of continuous outcomes, the sample mean and standard deviation (SD) are essential for synthesizing effect sizes across studies. However, clinical studies frequently report alternative summary statistics, such as the median, quartiles, and range. To enable inclusion of such studies, various methods have been proposed to estimate the sample mean and SD from these reported summaries. We propose the Bayesian Order Statistics-based Estimator (BOSE), which leverages the joint likelihood of observed order statistics together with weakly informative priors to obtain the full posterior distribution for the mean and SD without relying on computationally intensive iterative procedures such as Markov chain Monte Carlo algorithms. Our numerical studies demonstrate that BOSE performs competitively with existing approaches in estimating the mean, while achieving superior performance for estimating the SD across all evaluated scenarios, particularly in small-sample settings. Under non-normal distributions including skewed, heavy-tailed, and bimodal settings with mild or moderate deviations from normality, BOSE remains robust and stable, whereas methods specifically designed for skewed distributions may become unstable or even inapplicable. Beyond point estimation, BOSE naturally provides empirically validated posterior credible intervals, enabling researchers to formally quantify uncertainty for study-level estimates and make reliable, evidence-based decisions in meta-analytic research synthesis. A publicly accessible web application implementing BOSE and competing methods is also provided to facilitate practical use in meta-analytic research.
van de Beek, H.; Beldjenna, M.; Fidler, M. L.; Zwep, L. B.; van Hasselt, J. G. C.
Show abstract
Asymptotic standard errors for the parameters of a nonlinear mixed-effects model fitted by first-order conditional estimation (FOCE) or FOCE with interaction (FOCEI) require the observed (Fisher) information -- the negative second derivative of the population objective at the optimum. The gradient of this objective can be computed exactly from sensitivity equations, but the observed information is conventionally still formed by finite differencing, which is less accurate and step-size dependent. Our objectives are to (i) derive the FOCE and FOCEI observed information in closed form within the same sensitivity-equation framework, and (ii) quantify the precision this recovers. Writing the objective as a data term plus the log-determinant of the first-order inner Hessian, the population Hessian splits so that the data term reuses the second-order sensitivities already needed for the gradient, whereas the log-determinant term requires third-order sensitivity equations -- confining the third-order dependence to a single term, where it enters in exactly two places. A finite-difference error analysis shows the differenced Hessian attains an accuracy no better than the square root of the objectives evaluation accuracy, whereas the analytic form is limited only by the sensitivity and differential-equation solutions, with no step size to tune. We illustrate this on a one-compartment oral model with first-order absorption fitted to warfarin data, where the differenced standard error is usable only over a narrow band of step sizes while the analytic value carries none. Implemented in the open-source R package nlmixr2, the method makes exact, reproducible standard errors routine, supporting more dependable confidence intervals, identifiability assessment, and uncertainty propagation.